Papers with rule-based system

15 papers
Scalable and Robust Self-Learning for Skill Routing in Large-Scale Conversational AI Systems (2022.naacl-industry)

Copied to clipboard

Challenge: Existing methods to enable skill routing do not scale in terms of the number of skills and skill on-boarding.
Approach: They propose a model-based approach to enable natural conversation by allowing frequent policy updates . they propose an annotation-based system, rule-based model, and bandit-based learning .
Outcome: The proposed method is scalable and cost-effective, the authors show . they show that it can improve the user experience without abrupt policy changes .
A Simple Unsupervised Approach for Coreference Resolution using Rule-based Weak Supervision (2022.starsem-1)

Copied to clipboard

Challenge: state-of-the-art coreference models rely on labeled data, but an end-to-end model is needed to solve this problem.
Approach: They propose an approach that leverages an end-to-end neural model in settings where labeled data is unavailable.
Outcome: The proposed approach outperforms the previous best unsupervised model and outperformed the rule-based model on English OntoNotes corpus.
Team SVMrank: Leveraging Feature-rich Support Vector Machines for Ranking Explanations to Elementary Science Questions (D19-53)

Copied to clipboard

Challenge: TextGraphs 2019 Shared Task on Multi-Hop Inference for Explanation Regeneration tackles explanation generation for elementary science questions.
Approach: They propose a hybrid pipelined machine learning model and rule-based system to address MIER-19 . they use a featurerich learning-to-rank machine learning and a rule-driven system to rerank the LTR model predictions.
Outcome: The proposed model was ranked fourth in the evaluation, close to the second and third ranked teams, achieving 39.4% MAP.
Exploring Interpretability in Event Extraction: Multitask Learning of a Neural Event Classifier and an Explanation Decoder (2020.acl-srw)

Copied to clipboard

Challenge: EE is a key requirement for machine learning in many domains, e.g., legal, medical, finance.
Approach: They propose an interpretable approach for event extraction that jointly trains a classifier and a rule decoder for event processing.
Outcome: The proposed approach can be used for semi-supervised learning and its performance improves when trained on automatically-labeled data generated by a rule-based system.
Emotion Cause Extraction on Social Media without Human Annotation (2023.findings-acl)

Copied to clipboard

Challenge: Existing studies have focused on extracting emotion causes from news articles, but lack of fine-grained annotations has limited the ECE task.
Approach: They propose a new ECE framework that extracts emotion causes from social media data without relying on human annotations.
Outcome: The proposed framework achieves high extraction performance and generalizability without relying on human annotations.
Summarizing Patients’ Problems from Hospital Progress Notes Using Pre-trained Sequence-to-Sequence Models (2022.coling-1)

Copied to clipboard

Challenge: Problem list summarization requires a model to understand, abstract, and generate clinical documentation.
Approach: They propose a task that summarises patients' main problems from daily progress notes using input from the provider's progress notes during hospitalization.
Outcome: The proposed model outperforms two state-of-the-art seq2seq transformer architectures in summarizing patients' main problems from daily progress notes in the medical information mart for Intensive Care (MIMIC)-III.
Enhancing Sequence-to-Sequence Neural Lemmatization with External Resources (2021.eacl-main)

Copied to clipboard

Challenge: a hybrid approach to lemmatization enhances the seq2seq neural model with additional lemmas extracted from an external lexicon or a rule-based system.
Approach: They propose a hybrid approach that enhances a seq2seq neural model with additional lemmas extracted from an external lexicon or a rule-based system.
Outcome: The proposed model achieves statistically significant improvements on 23 UD languages, compared to baseline models not utilizing additional lemma information.
HarfoSokhan: A Comprehensive Parallel Dataset for Transitions between Persian Colloquial and Formal Variations (2026.eacl-long)

Copied to clipboard

Challenge: A wide array of NLP/NLU models have been developed for the Persian language but performance drops when applied to the colloquial form of Persian.
Approach: They propose to use a large-scale colloquial to formal Persian parallel dataset to train a GPT2 model that exhibited remarkable proficiency in colloqual to informal text style transfer.
Outcome: The proposed dataset outperforms OpenAI’s GPT-3.5-turbo model and a leading rule-based system in colloquial to formal Persian conversion.
Harnessing Pre-Trained Neural Networks with Rules for Formality Style Transfer (D19-1)

Copied to clipboard

Challenge: Existing studies normalize informal sentences with rules, but they introduce noise if we use them in a naive way.
Approach: They propose to harness rules into a state-of-the-art neural network that is typically pretrained on massive corpora.
Outcome: The proposed method can be used to generate a state-of-the-art on a small dataset.
Intrinsic Task-based Evaluation for Referring Expression Generation (2024.acl-long)

Copied to clipboard

Challenge: Referring Expression Generation (REG) models generate referring expressions that refer to referents at different points in a discourse.
Approach: They propose to use a purely ratings-based human evaluation to evaluate REG models by completing two meta-level tasks.
Outcome: The proposed evaluation makes the models more reliable and discriminable, and improves the quality of the REs.
GeNRe: A French Gender-Neutral Rewriting System Using Collective Nouns (2025.findings-acl)

Copied to clipboard

Challenge: Gender rewriting is an NLP task that uses gendered forms to mitigate gender biases.
Approach: They propose a French gender-neutral rewriting system using collective nouns, which are gender-fixed in French.
Outcome: The proposed system detects gendered forms and replaces them with neutral or opposite forms.
The Use of Text Alignment in Semi-Automatic Error Analysis: Use Case in the Development of the Corpus of the Latvian Language Learners (L18-1)

Copied to clipboard

Challenge: Using error annotation methods, the corpus of the Latvian language learners can be adapted for other languages with relatively free word order.
Approach: They propose a method for creating error annotated corpora using text correction, automated morphological analysis, automated text alignment and error annotation.
Outcome: The proposed method has been approbated in the development of the corpus of the Latvian language learners.
Replace and Report: NLP Assisted Radiology Report Generation (2023.findings-acl)

Copied to clipboard

Challenge: Clinical practice frequently uses medical imaging for diagnosis and treatment.
Approach: They propose a template-based approach to generate radiology reports from radiographs . they use multilabel image classifiers to generate tags, pathological descriptions from tags .
Outcome: The proposed method improves on the most popular radiology report datasets.
Enhancing Accessible Communication: from European Portuguese to Portuguese Sign Language (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing systems for translating European Portuguese into LGP glosses rely on hand-crafted rules . current systems rely only on toy examples, disregarding non-manual movements .
Approach: They propose a corpora-driven rule-based machine translation system between European Portuguese and LGP glosses and two neural machine translation models.
Outcome: The proposed system improves on existing translation systems and annotates a gold collection of the results.
PictoEduca: Building a Dataset for Spanish Text-to-Pictogram Generation (2026.findings-acl)

Copied to clipboard

Challenge: PictoEduca is the first large-scale Spanish text-to-pictogram dataset for augmentative and alternative communication.
Approach: They present PictoEduca, a large-scale Spanish text-to-pictogram dataset for augmentative and alternative communication.
Outcome: The proposed dataset combines automatic annotation with targeted expert correction, supporting scalable and high-quality corpus construction.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations